Processing Natural Malay Texts: a Data-driven Approach

نویسنده

  • Zuraidah Mohd Don
چکیده

This research represents the first attempt to produce a working system for the automatic processing of texts of Bahasa Melayu ‘Malay’. At the heart of the system is an integrated relational lexical database called MALEX, which draws on the experience of working on English and other languages, but which is specifically tailored to the conditions of Malay. The development of the database is from the beginning entirely data driven, and is based on the analysis of a corpus of naturally produced Malay texts. In designing procedures which access the database, properties of the text are consistently and rigorously distinguished from properties of the lexicon and of the grammar. The system is currently used to provide information for a range of applications, for grammatical tagging, stemming and lemmatisation, parsing, and for generating phonological representations. It is hoped and intended that the design features of MALEX will be transferable, and provide a model for the development of working systems for other Asian languages.

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

An Exploratory Study of the Malay Text Processing Tools in Ontology Learning

This paper discusses the overall process of learning taxonomy from Malay texts using unsupervised conceptual clustering approach and investigates the existing Malay NLP tools as potential pre-processing tools for the proposed ontology learning approach. The tools are a maximum-entropy parser based on open NLP package, a word sense tagger and a parser based on pola grammar. A case study approach...

متن کامل

An architecture for Malay Tweet normalization

Research in natural language processing has increasingly focused on normalizing Twitter messages. Currently, while different well-defined approaches have been proposed for the English language, the problem remains far from being solved for other languages, such as Malay. Thus, in this paper, we propose an approach to normalize the Malay Twitter messages based on corpus-driven analysis. An archi...

متن کامل

Computational Linguistics at Universiti Sains Malaysia

This paper gives a brief history of UTMK, a computer-aided translation unit, and reports on her projects and research co-operations. After its beginnings as a thesis project on Malay affixation, UTMK’s interest moved from machine translation to the development of tools for translation. Today, UTMK’s focus is on the development of natural language processing applications and tools (internet brow...

متن کامل

Mining Opinion in Online Messages

The number of messages that can be mined from online entries increases as the number of online application users increases. In Malaysia, online messages are written in mixed languages known as ‘Bahasa Rojak’. Therefore, mining opinion using natural language processing activities is difficult. This study introduces a Malay Mixed Text Normalization Approach (MyTNA) and a feature selection techniq...

متن کامل

De-Constraining Text Generation

We argue that the current, predominantly task-oriented, approaz~hes to modularizing text • generation, while plausible and useful conceptually, set up spurious conceptual and operational constraints. We propose a data-driven approach to modularization and illustrate how it eliminates • •the previously ubiquitous constraints on combination of evidence across modules and on • control. We also bri...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

عنوان ژورنال:

دوره   شماره 

صفحات  -

تاریخ انتشار 2010